Back

BioData Mining

Springer Science and Business Media LLC

Preprints posted in the last 30 days, ranked by how well they match BioData Mining's content profile, based on 22 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.

1
A Semantic + Neuronal Approach to Predict Pathogenic Variants in DNA Sequences

Motta, J. A.; Motta, M. d. M.; Fernandez, C.

2026-08-20 bioinformatics 10.64898/2026.08.16.745093 medRxiv
Top 0.1%
14.8%
Show abstract

In this work, we present a machine learning model for identifying pathogenic DNA variants. The model was learned from the analysis of normal and pathogenic sequences extracted from the ClinVar database (supported by NCBI). This analysis was based on a conceptual semantic model of DNA sequences converted to peptide sequences (amino acid sequences) governed by a well-defined grammar, which allowed us to apply NLP techniques, specifically Part of Speech tagging (POS tagging). Our predictive model was built by combining two techniques: CRF (from the Markov model family), which performs the sequencing, and BiLSTM (a deep learning model) which captures the past and future content of the sequences. The training space was created with the sequences of 105 genes associated with approximately 27,000 pathogenic variants. The model was evaluated using the metrics precision, P-R and ROC curves, AUC, and confusion matrices. Its performance was also compared against five known methods for predicting pathogenic variants. The results show exceptional performance that exceeds expectations and places this new method at the state of the art for predicting pathogenic DNA sequences.

2
Beyond Chemical Similarity: Structure-Agnostic Drug-Drug Interaction Prediction with MeSH Semantics and a Drug-Target-Protein Knowledge Graph

Yılmaz, A.; Szydlik, S.; Taheri, G.

2026-08-18 bioinformatics 10.64898/2026.08.10.743843 medRxiv
Top 0.1%
8.0%
Show abstract

BackgroundAdverse drug-drug interactions (DDIs) cause preventable hospitalizations, but exhaustive experimental screening of all drug pairs is infeasible. Many computational predictors rely on SMILES or other molecular representations, limiting their direct applicability to biologics and other non-small-molecule therapeutics. We present a structure-agnostic framework that combines semantic representations derived from Medical Subject Headings (MeSH) with graph-derived topology from a Drug-Target-Protein knowledge graph constructed from DrugBank and UniProt. We further investigate how variation in MeSH annotation depth affects predictive performance. ResultsDrugs are grouped according to their deepest MeSH annotation level (Low, Mid, or Deep), and performance is evaluated across the resulting interaction categories in transductive and inductive settings. The Intermediate ontology scope (Low+Mid) provides the most stable performance, while adding Deep-level terms offers limited and inconsistent benefit. Lightweight topological descriptors are integrated with MeSH features through instance-wise, dimension-specific latent-space gating, using curated reliable-negative pairs for supervision. Fusion improves mean performance over the MeSH-only baseline across all six categories in the transductive setting. Under induction, the clearest gains occur for Low-Low interactions ({Delta}AUROC = 0.056;{Delta} F1 = 0.137) and Low-Mid interactions ({Delta}AUROC = 0.077;{Delta} F1 = 0.114). ConclusionsMeSH annotation depth is associated with systematic variation in DDI prediction performance that aggregate evaluation can obscure. Graph-derived topology is particularly beneficial when ontology annotations are shallow. The framework provides a common, structure-agnostic representation compatible with both small-molecule and biologic therapeutics and supports first-pass DDI prioritization for subsequent expert assessment.

3
Automatic bioinformatic software named entity recognition from literature

Xuan, H.; Pasupuleti, R.; Liu, B.; Sun, H.; Zhang, J.; Yao, Z.; Zhong, C.

2026-09-01 bioinformatics 10.64898/2026.08.26.731133 medRxiv
Top 0.1%
6.2%
Show abstract

Bioinformatics software and databases are essential components of modern life science research, yet their mentions in the scientific literature are often inconsistent and difficult to systematically identify at scale. The lack of a comprehensive and up-to-date catalog of bioinformatics resources hinders efforts toward automated biomedical knowledge extraction and streamlined data analysis. Here we present SNAIL, a hybrid named entity recognition framework designed to automatically identify bioinformatics software and database (SW/DB) names from biomedical texts. SNAIL integrates complementary lexical and semantic modeling strategies. The lexical component captures orthographic patterns and contextual cues characteristic of SW/DB names, while the semantic component leverages contextual embeddings generated by transformer-based language models such as SciBERT, combined with an explicit token-masking strategy to enhance entity-focused representations. A large training corpus was constructed automatically through a hybrid pipeline that integrates citation-hinted extraction with large language model-assisted distillation. Evaluation on two independent benchmark datasets and real-world research articles demonstrates that SNAIL substantially outperforms existing approaches, including domain-specific methods such as bioNerDS2 and general-purpose large language models such as ChatGPT, Gemini, Grok and Claude. Applying SNAIL to large-scale literature analysis further reveals distinct journal-level preferences across bioinformatics subfields. These results demonstrate that SNAIL provides an accurate and scalable solution for identifying bioinformatics resources in scientific texts and enables systematic meta-analysis of tool usage and research trends.

4
Uncovering High-Order Epistatic Interactions in GWAS via a Machine Learning-Based Feature Engineering Framework

Byun, J.; Saha, D.; Han, Y.; Shaw, V. R.; Siminovitch, K.; Amos, C. I.

2026-08-09 genomics 10.64898/2026.08.03.742638 medRxiv
Top 0.1%
6.2%
Show abstract

BackgroundGenome-wide association studies (GWAS) often fail to identify higher-order epistatic interactions that contribute to complex inheritance patterns of traits and diseases. While machine learning (ML) can capture non-linear relationships, extracting interpretable insights from these models remains a challenge. We propose a novel tree-based feature engineering framework that uses Classification and Regression Trees (CART) to explicitly encode high-order interaction decision paths as dummy variables. We investigate three path-based encoding strategies: (i) all decision paths, (ii) leaf-node paths only, and (iii) internal-node paths only. This approach aims to transform complex decision boundaries into discrete features that capture nonlinear interactions that are not readily captured by traditional association models. ResultsThe framework was evaluated using genetic data for ANCA-associated vasculitis (AAV). To manage the high dimensionality of the engineered feature space, we applied a comprehensive suite of ML methods across three tasks: (1) Ensemble Learning (Random Forest, XGBoost, and Gradient Boosting Machine); (2) Decision Tree Analysis (CART); and (3) Regression and Classification Tasks (Regularized Linear Regression/LASSO, Support Vector Machine, and Logistic Regression). Stepwise feature selection and regularization were employed to isolate the most informative interaction patterns. Results indicate that incorporating CART-derived interaction paths--particularly those from high-impact regions of the tree--significantly improves classification accuracy and model interpretability compared to using the original feature space alone. ConclusionsThe proposed framework provides a robust, scalable methodology for identifying high-order genetic interactions. By bridging the gap between the predictive power of ensemble ML and the necessity for mechanistic insight, this approach offers a clearer mapping of the combinatorial genetic processes underlying complex diseases. While applied here to AAV, the method is highly adaptable for exploring the genetic architecture of diverse populations and complex traits.

5
Electronic health data exploring cardiorespiratory responses of transfusions in preterm infants: An international multicenter cohort study

Honore, A.; Rech, T.; Scrivens, A.; Binotto, I.; Zandvoort, C. S.; van der Staaij, H.; Peck, M.; Zivanovic, S.; Stanworth, S. J.; Hartley, C.; Dame, C.; Deschmann, E.; the Neonatal Transfusion Network,

2026-09-03 pediatrics 10.64898/2026.09.01.26361418 medRxiv
Top 0.1%
5.6%
Show abstract

Background and Objectives: Preterm infants are commonly transfused, yet direct cardiorespiratory effects of red blood cell (RBC) transfusions remain poorly understood. We explored the feasibility of using multicentre electronic health data (EHD) to study such cardiorespiratory responses. Methods: Highly granular routine EHD were collected from preterm infants born <32 weeks gestational age at three European centres. Heart rate, oxygen saturation, and respiratory rate were evaluated 12 hours before and after the RBC transfusion. Results: A total of 321 transfusions in 164 infants were analysed. Overall, there was no significant change in the rate of bradycardia and apnoea following transfusion. Cardiorespiratory parameters varied substantially between infants; e.g. 20% of transfusions were associated with an unexpected, significant increase in heart rate. Respiratory rate and oxygen saturation exhibited similarly heterogenous patterns following transfusion. In sub-group analysis, the proportion of transfusions with increased heart rate was significantly higher within the first two weeks than later (32% vs 13%, p=0.0019). Conclusions: Multicentre EHD extraction allows to identify otherwise masked short-term effects of RBC transfusions on cardiorespiratory parameters, possibly indicating cardiac or pulmonary overload. Such effects may vary with adaptation to anaemia. Analysing EHD may ultimately enable personalized transfusion practice.

6
Identifying multi-omics biomarkers for ovarian cancer survival estimation

Fateh, K.; Yerukala Sathipati, S.

2026-08-10 bioinformatics 10.64898/2026.08.04.742866 medRxiv
Top 0.1%
4.9%
Show abstract

Ovarian cancer is among the deadliest gynecologic malignancies, and its molecular heterogeneity limits accurate prognostic stratification. Although multi-omics approaches have improved predictive modeling, many prioritize predictive performance over biological interpretability, limiting their clinical translation. We developed an interpretable three-stage machine learning framework integrating mRNA, microRNA, DNA methylation, copy number variation, and protein expression data from The Cancer Genome Atlas. Hierarchical feature selection was combined with a weighted ensemble of ElasticNet, ridge regression, support vector regression, XGBoost, and random forest models to estimate overall survival time in patients with ovarian cancer. Multi-omics integration outperformed every single-modality model, achieving a Pearson correlation of 0.752, a concordance index of 0.779, and a mean absolute error of 8.57 months between estimated and observed survival time, compared with 0.48 for the best single modality. The framework identified a 20-biomarker signature dominated by tumor-associated macrophage and complement genes. In an independent survival analysis, VSIG4 and CD163 remained significant after false discovery rate correction, and the signature raised the concordance index over clinical covariates alone from 0.615 to 0.686Enrichment analysis implicated PI3K-Akt, MAPK, focal adhesion, hypoxia, apoptosis, and p53 signaling pathways. This framework couples improved prognostic estimation with biological interpretability supporting multi-omics biomarker discovery in ovarian cancer.

7
Assessing Computational Models for Pharmacogenomic Variant Interpretation

Pucci, F.; Hermans, P.; Tsishyn, M.; Cusato, J.; Rooman, M.

2026-08-09 bioinformatics 10.64898/2026.08.03.742561 medRxiv
Top 0.1%
4.8%
Show abstract

Accurately predicting the effects of pharmacogenomic variants is essential for the development of personalized therapeutic strategies, as genetic variability can influence drug response differently across patients. Here, we assessed several computational approaches using a dataset of pharmacogenomic variants with either clinical annotations or functional characterization by deep mutational scanning, compiled from the literature, with an additional focus on CYP2C9, a clinically relevant drug-metabolizing enzyme. Our results show that, despite recent methodological advances, substantial room for improvement remains. In particular, current methods struggle to distinguish gain-of-function variants associated with increased drug clearance and fast-metabolizer phenotypes from neutral variants, whereas loss-of-function variants that reduce drug clearance are predicted more accurately. The integration of structural and evolutionary information appears to be a key strategy for improving performance, with the coevolution-based StructureDCA method achieving the highest accuracy compared with classical genetic variant-effect predictors and recent deep learning approaches, including the pathogenic-variant predictor AlphaMissense and general protein language model-based methods. Finally, our results indicate that computational models can complement in vitro experiments in clinical variant interpretation, as StructureDCA predictions showed better agreement with clinically annotated phenotypes than large-scale deep mutational scanning data in several cases.

8
Feasibility of adjusting for sepsis-related organ dysfunction in pediatric patients using administrative healthcare data

Ravichandrajah, H.; Fischer, A.; Tiago Gomez, A.; Hojeij, R.; Goretzki, S. C.; Felderhoff-Mueser, U.; Park, H.-J.; Kernan, K.; Carcillo, J. A.; Dohna-Schwake, C.; Bruns, N.

2026-08-13 pediatrics 10.64898/2026.08.12.26360255 medRxiv
Top 0.1%
4.3%
Show abstract

Background: Risk adjustment for disease severity in pediatric intensive care research commonly relies on clinical organ dysfunction scores requiring detailed clinical and laboratory information, which is often unavailable in administrative healthcare datasets. We therefore evaluated the feasibility of a coding-based Pediatric Organ Dysfunction Index (PODI) derived from International Classification of Diseases (ICD-10) and Operation and Procedure System (OPS) codes, for approximating sepsis-related organ dysfunction and adjusting for disease severity, using the pediatric Sequential Organ Failure Assessment (pSOFA) score as a reference standard. Methods: In this retrospective single-center cohort study, pediatric sepsis episodes treated between November 2011 and November 2021 were identified. Discrimination for in-hospital mortality and calibration were assessed. Agreement between PODI and pSOFA was quantified using Spearman's rank correlation, and organ-specific agreement using sensitivity, specificity, and predictive values. An expanded PODI incorporating additional ICD-10 and OPS codes was evaluated in sensitivity analyses. Results: A total of 488 pediatric sepsis episodes were included, with an in-hospital mortality of 14.1%. The PODI showed good discrimination for in-hospital mortality (AUC 0.85, 95% CI 0.80-0.89), comparable to the maximum pSOFA (pSOFAmax) (AUC 0.78, 95% CI 0.72-0.83) and superior to pSOFA at sepsis onset (pSOFAonset) (AUC 0.73, 95% CI 0.67-0.80). Agreement between PODI and pSOFA organ-specific components varied considerably across organ systems, with the highest sensitivity to detect pulmonary dysfunction. Correlation between both scores was moderate (0.54 for pSOFAonset and 0.60 for pSOFAmax), indicating that comparable predictive performance does not render the scores interchangeable. The expanded PODI improved organ-level sensitivity for selected components but did not meaningfully improve mortality discrimination. Conclusions: The standard PODI may represent a practical approach to adjust for organ dysfunction and therapy intensity in administrative datasets with ICD-10 coding where clinical and laboratory information is unavailable. Given only moderate agreement with the pSOFA, the PODI should be understood as a covariate for risk adjustment at the group level rather than as a substitute for clinical organ dysfunction scores in individual patients. Further validation and refinement in non-sepsis cohorts are required before broader implementation in large-scale administrative research can be recommended.

9
Mechanistic Multi-Task Logistic Regression as an Alternative to Parametric Hazard Models in Joint Time-to-Event Analysis

Bisaso, K. R.; Kadada, K. R.; Bisaso, K. S.; Ette, E. I.

2026-08-18 pharmacology and therapeutics 10.64898/2026.08.15.26360512 medRxiv
Top 0.1%
4.1%
Show abstract

Background: Parametric time-to-event models require specification of a baseline hazard function, which may influence prediction when the underlying hazard shape is uncertain. This study compared conventional joint longitudinal time-to-event models with mechanistic Multi-Task Logistic Regression, which directly models the survival distribution without selecting a continuous parametric hazard family. Methods: A simulated dataset of 100 individuals with longitudinal sum of longest diameters and event outcomes was analyzed using a shared mechanistic tumor shrinkage regrowth model. Event submodels comprised exponential, Gompertz, Weibull, log-normal, log-logistic, and circadian hazards, mechanistic Multi-Task Logistic Regression, and a hybrid neural-mechanistic extension. All models were estimated jointly using shared patient-specific random effects and longitudinal data. Models were evaluated using longitudinal goodness-of-fit, visual predictive checks, five-fold cross-validated inverse-probability-of-censoring-weighted dynamic area under the curve and Brier scores, integrated Brier score, calibration, and event-interval negative log score. Results: Longitudinal parameter estimates and diagnostics were comparable across models. All conventional hazard models produced identical dynamic area under the curve values within prediction windows, although probabilistic accuracy differed. The log-normal hazard achieved the lowest overall integrated Brier score (0.1928). Mechanistic Multi-Task Logistic Regression achieved the highest later landmark discrimination (area under the curve 0.867 versus 0.798 for all hazard models) and the lowest mean event-interval negative log score (2.362). The hybrid model improved intermediate-landmark discrimination but not overall probabilistic accuracy. Conclusions: Mechanistic Multi-Task Logistic Regression provided competitive joint time-to-event prediction while avoiding baseline hazard-family selection. It represents a practical complementary alternative to parametric hazard modeling, particularly when hazard shape is uncertain and dynamic discrimination is important.

10
Detecting CYP2C19 deletions from genotyping array signals using neural networks

Yelmen, B.; Hofmeister, R. J.; Lutsar, V. K.; Finianos, M.; Stone, B. C.; Joeloo, M.; Krebs, K.; Kivistik, P. A.; Smit, S.; Estonian Biobank Research Team, ; Metspalu, M.; Hudjashov, G.; Milani, L.

2026-08-25 bioinformatics 10.64898/2026.08.21.746170 medRxiv
Top 0.1%
4.1%
Show abstract

Since copy number variations (CNVs) in pharmacogenes can cause significant alterations in drug metabolism, their reliable detection is of high importance both for large-scale studies and personalized medicine. Whole-genome sequencing, and specifically long-read sequencing, is the gold standard for CNV detection. Despite increasing availability of these technologies, genotyping arrays are still widely used as cost-effective alternatives in biobank and clinical settings, yet calling CNVs based on array intensity signals is challenging due to low base pair resolution. In this work, we developed a neural network model, nnCNV, to predict deletions in the CYP2C19 pharmacogene region from array intensity signals. We compared our method to the most widely used algorithm, PennCNV, and demonstrated better performance reaching 100% accuracy in the test dataset. Furthermore, we predicted probe-by-probe CYP2C19 deletion coordinates for all Estonian Biobank samples using nnCNV and PennCNV, and validated these predictions using an identity-by-descent (IBD) sharing method, which also demonstrated superior nnCNV performance. For the deletion samples with conflicting PennCNV and nnCNV predictions, we performed PCR analysis for validation, which showed 97% precision for nnCNV compared to 23% for PennCNV. Finally, we assessed the gradient-based feature importance maps and showed that nnCNV utilizes signal intensity information not only from deletion probes, but also from probes in flanking regions. Our results demonstrate that long-range information, which cannot be utilized by hidden Markov models, can improve CNV calling.

11
Evaluating Aggregated Gene Level eQTL Scores

Meyer, D.; Popko, N.; Laub, D.; Schofield, P.; Amariuta, T.; Alexandrov, L. B.; Carter, H.

2026-08-26 bioinformatics 10.64898/2026.08.21.746287 medRxiv
Top 0.1%
4.0%
Show abstract

Genetic feature engineering, used in methods such as transcriptome-wide association study, supports gene-trait association testing by aggregating single variants into gene-level features predictive of expression. To evaluate how different model architectures, LD filtering thresholds, and variant prioritization methods affect expression prediction quality, we trained over 3 million models and evaluated their performance in independent cohorts. Using the best performing models to impute expression and immunotherapy response as an example trait, we found a significant association with the reactive oxygen species pathway (p=0.032). Our model training workflow will support genetic feature engineering towards improved complex trait modeling.

12
Global protein expression profiling in stem cell factor stimulated human Acute megakaryoblastic leukemia cells identifies CFL1, GSN and CCT8 as prognostic biomarkers for Acute Myeloid Leukemia.

Ravi, A. K.; Gopan, G.; Arumugam, S.; Sethumadhavan, A.; Mani, M.

2026-08-26 cancer biology 10.64898/2026.08.24.746695 medRxiv
Top 0.1%
3.5%
Show abstract

Abstract Background: The stem cell factor receptor or c-Kit is a type III receptor tyrosine kinase, activated by its ligand Stem cell factor (SCF). Up on activation, c-kit induces signaling pathways that regulates blood cell proliferation, survival, differentiation, and migration. Several studies reported that c-Kit/SCF signaling, contributes to the development and progression of acute myeloid leukemia (AML) in patients. However, the downstream proteins regulated by c-kit activation and their clinical significance in AML remain poorly explored. Methods: Human Acute megakaryoblastic leukemia (Mo7e) cells, were-stimulated with SCF and global protein expression were profiled using two-dimensional gel electrophoresis coupled with MALDI-TOF and LC-MS/MS. Differentially expressed proteins were functionally characterized and validated using patient data from the TCGA-LAML and matched normal data from GTEx, GEO datasets, and quantitative RT-PCR. Their diagnostic and prognostic significance was assessed using ROC, Cox regression, LASSO, Kaplan Meier survival analyses, and a prognostic nomogram model. Results: Proteomic profiling identified 14 differentially expressed proteins in SCF-stimulated Mo7e cells, which are predicted to involved in cytoskeletal organization, protein folding, metabolism, vesicular trafficking, and translational regulation. Transcriptomic analysis of the TCGA-LAML cohort revealed significant dysregulation of CFL1, CCT8, HSP90B1, MDH2, EIF5A, GSN, and TPI1. Integrated ROC, Cox regression, and LASSO analyses identified CFL1, CCT8, and GSN as the most robust prognostic biomarkers associated with poor overall survival in LAML patients. Their expression patterns were validated in independent GEO datasets and by qRT-PCR in SCF stimulated Mo7e cells. Finally, a three-gene nomogram model was developed and validated to predict the overall survival probability of AML patients at 1-, 3-, and 5-year time points. Conclusions: This study identifies CFL1, CCT8, and GSN as key downstream effectors of c-Kit signaling as prognostic biomarkers for AML. These findings provide mechanistic insights into c-Kit-driven leukemogenesis and establish a clinically relevant three-gene signature for AML risk stratification and potential therapeutic targeting.

13
Bayesian Borrowing of External Information in Clinical Trials: A Comparison of MAP, RMAP, and SAM Priors

Choi, L.; McNeer, E.; Beck, C. A.; Neul, J. L.

2026-08-31 pharmacology and therapeutics 10.64898/2026.08.26.26360843 medRxiv
Top 0.1%
3.5%
Show abstract

Bayesian borrowing of external information can improve trial efficiency, particularly in pediatric and rare disease settings where patient populations are limited, but may introduce bias and inflate the Type~I error rate when the trial differs from external studies. Recent U.S. Food and Drug Administration (FDA) draft Bayesian guidance emphasizes careful evaluation of external information, prior specification, and assessment of operating characteristics. This paper compares three meta-analytic-predictive (MAP)-based methods for Bayesian borrowing: the MAP prior, robust MAP (RMAP) prior, and self-adapting mixture (SAM) prior. An adaptive platform trial design in Rett syndrome is used as a case study. Simulation studies evaluate frequentist operating characteristics under varying prior--data conflict, between-study heterogeneity, treatment effects, and clinically significant differences (CSDs) for the SAM prior. The MAP prior achieved the greatest efficiency when external and current data were compatible but exhibited the largest bias under substantial prior--data conflict. The RMAP priors improved robustness through fixed robust-component weights, whereas the SAM prior adaptively adjusted borrowing and was less sensitive to prior--data conflict while retaining efficiency gains when the data were compatible. Although the CSD influenced the degree of adaptive borrowing, as reflected by effective sample size, it had only a modest impact on frequentist operating characteristics. Sensitivity analyses using a skeptical robust component yielded similar qualitative conclusions, while accentuating the differences between the MAP and RMAP priors. These findings provide guidance for evaluating and selecting MAP-based borrowing strategies before trial implementation, particularly in rare disease settings, consistent with current FDA recommendations.

14
A Guided AI Framework for Customizable and Efficient Harmonisation to the OMOP Common Data Model

Nehra, N.; Swami, R.; Dadi, D.; Mishra, R.; Sharma, U.; Verma, P.; Sen, M.; Dhruw, N. K.; Jha, A. K.

2026-08-12 bioinformatics 10.64898/2026.08.07.742453 medRxiv
Top 0.1%
3.4%
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWGetting clinical data from different sources to "talk" to each other within the OMOP Common Data Model (CDM) is arguably the most tedious part of multi-center research. While this integration is essential, the transformation process is frequently a manual grind, requiring a rare overlap of deep clinical knowledge and technical expertise. In this paper, we present a framework designed to alleviate some of the burden on the researcher by automating data harmonization through two distinct steps: structural schema mapping and terminological standardization. For the structural piece, we moved away from "black box" logic in favor of a stateful workflow managed by large language models (LLMs) and directed acyclic graphs. By profiling EHR data at the source, our system generates context-aware dictionaries that offer ranked mapping suggestions alongside confidence scores. While our benchmarking showed a 97.5% agreement rate at the schema level and an 84% agreement rate at the value level when compared with human experts, the system appears most effective when treated as a "co-pilot" rather than a total replacement for human oversight. To handle value-level standardization, we implemented a hybrid search strategy that pairs the semantic depth of SapBERT embeddings with the literal precision of fuzzy string matching. By using FAISS for rapid similarity retrieval, the engine attempts to resolve messy or "noisy" clinical descriptions to standard OMOP concepts. This approach seems particularly promising for handling the non-standardized labels that often plague smaller, local datasets. Ultimately, our results suggest that this guided approach can shift the timeline for OHDSI-compliant warehousing from weeks of manual curation to a more manageable and scalable pipeline, potentially lowering the barrier to entry for smaller research teams.

15
BfBio: a graph-based tool for the prediction of Angiogenic Stalk Cell genes using a Personalized PageRank algorithm

Bettoni, L.; Dmitrieva, J.; Mousa, M.; Alsafar, H.; Saeys, Y.; Zakeri, P.; Carmeliet, P.

2026-08-24 cancer biology 10.64898/2026.08.23.746493 medRxiv
Top 0.1%
3.2%
Show abstract

Although most human protein coding genes have functional annotations in databases, such as GeneCards, many remain poorly characterized. To address this gap, computational tools can be leveraged to predict the functional roles of under-annotated genes by extracting patterns from complex biological networks. Here we introduce Brain-for-Biotech (BfBio), a framework designed to identify genes important for vascular endothelial cells (EC), which are crucial cells for vessel formation (angiogenesis), vascular homeostasis, hemostasis and blood/tissue barrier function but also critical mediators of immunity and cancer progression. BfBio utilizes a Personalized PageRank (PPR) algorithm on an integrated network of different omics datasets and publicly available gene-gene/protein-protein interaction databases. In this study, we apply the predictive capabilities of BfBio to infer angiogenic stalk cell phenotype function in genes for which this function was not known before. By leveraging a set of genes characterizing the stalk cell cluster in lung tumor EC models previously identified, we have achieved a high Area Under Receiver Operative Characteristic (AUC-ROC) performance (0.837). Enrichment analysis, coupled with a text mining application, further confirmed that among the 49 predicted genes four of them were poorly characterized yet possessed biologically relevant properties and were linked to cancer, thereby validating BfBio as a robust tool for prioritizing novel therapeutic targets in vascular biology.

16
Assessing the Reliability of LLM-Generated Phenotype-Genotype Associations Through External Validation

Sun, C.; Xin, Y.; Zeng, S.; Sunthankar, S. D.; Su, W.-C.; Lynn, J.; Mundo, S.; Babanejad, M.; Feng, Q.; Wei, W.-Q.

2026-08-21 bioinformatics 10.64898/2026.08.13.744701 medRxiv
Top 0.1%
3.1%
Show abstract

BackgroundPhenotype-genotype associations underpin precision medicine by enabling disease prevention, early diagnosis, risk stratification, therapeutic target discovery, and personalized treatment. However, the rapid growth of scientific evidence has made manual curation of these associations increasingly labor-intensive, time-consuming, and incomplete. Large Language Models (LLMs) offer a potential path to scalable genomic generation and synthesis of this knowledge, but their ability to accurately identify phenotype-genotype associations and the extent to which these outputs are supported by established genomic knowledge bases remain unclear. Materials and MethodsFour LLMs, Claude Sonnet 4.6, DeepSeek V4 Flash, Gemini 3 Flash Preview, and GPT-5.5, were benchmarked on six zero-shot task categories covering forward and reverse phenotype-gene and phenotype-SNP generation. A total of 4,196 associations were identified from curated inputs and evaluated through a multistage external verification pipeline comprising phenotype normalization, ontology mapping, genomic identifier validation against Ensembl, and evidence verification using both the GWAS Catalog and OMIM. Associations were assigned a fused evidence level of strong, moderate, weak, or none. ResultsOverall, 74.19% of generated associations were matched to at least one external genomic knowledge base; 9.15% received strong support and 54.46% moderate support. Phenotype-gene associations were more verifiable than phenotype-SNP associations (strong or moderate: 67.19% vs 54.06%). Among existing associations, Claude Sonnet 4.6 achieved the highest overall strong or moderate rate (69.2%), followed by GPT-5.5 (65.1%), DeepSeek V4 Flash (61.7%), and Gemini 3 Flash Preview (56.9%). ConclusionLLMs can support scalable generation of candidate phenotype-genotype associations. Performance varied substantially by relation type and was lower for SNP-level and rare disease associations, highlighting both the limitations of current genomic resources and the need for rigorous validation pipelines.

17
An M-learner approach for heterogeneous mediation analysis with high-dimensional omics mediators

Li, X.; Wei, P.

2026-09-01 bioinformatics 10.64898/2026.08.25.747106 medRxiv
Top 0.1%
3.1%
Show abstract

Causal mediation analysis is widely used to identify biological pathways linking exposures to outcomes, but most methods assume homogeneous mediation effects across individuals. In high-dimensional omics settings, this assumption can mask important heterogeneity driven by demographic, genetic, or environmental factors. We propose the M-high-learner, a flexible framework for detecting heterogeneous mediation effects with high-dimensional mediators. The method identifies mediators with subgroup-specific indirect effects while distinguishing them from null or homogeneous signals and controlling the type I error rate. It is computationally efficient, scalable, and yields interpretable sub-types. Simulation studies show that the proposed approach achieves high power while maintaining accurate error control. Applications to the Framingham Heart Study and the Multi-Ethnic Study of Atherosclerosis reveal that the mediation role of gene expression in sexs effect on high-density lipoprotein varies across subgroups defined by body mass index and age. Our framework provides a practical tool for uncovering heterogeneous biological mechanisms in high-dimensional genomic studies. Author SummaryBiological processes linking risk factors to disease often differ across individuals, but many existing methods assume these processes are the same for everyone. This can hide important differences between groups. We developed a powerful method to identify when these pathways vary across subgroups using large-scale molecular data. Our approach detects differences in how intermediate biological factors contribute to outcomes in populations defined by characteristics such as age and body mass index. Applying our method to population studies, we found that some biological pathways operate differently across groups, suggesting that key mechanisms may be missed when differences are ignored. Our work provides a tool to better understand how disease-related processes vary across individuals, which may support more targeted and personalized approaches to health research.

18
Model Validation Protocols for Machine Learning in Small Molecule Drug Discovery

Seal, S.; Zalte, A. S.; Araripe, D. A.; Gomes, R. A.; Korani, D.; Shekhar, M.; Siramshetty, V. B.; Patra, A.; Mou, Z.; Yu, X.; Kuhn, D.; Weskamp, N.; Ash, J.; Cheng, A. C.; Fang, C.; Price, D.; Aldeghi, M.; Rodriguez-Perez, R.; Clevert, D.-A.; Engkvist, O.; Deibler, K.; Rouquie, D.; Reutlinger, M.; Richmond, N. J.; Ainsley, J.; Ledeboer, M.; Green, W. H.; Bender, A.; Wognum, C.

2026-08-24 bioinformatics 10.64898/2026.08.19.745868 medRxiv
Top 0.1%
2.8%
Show abstract

Machine learning (ML) models for molecular property prediction are increasingly deployed in drug discovery, yet their adoption in real-world scenarios requires an understanding of the conditions in which a model succeeds or fails. While standardized benchmarks are powerful instruments to measure and unlock progress in ML research, they should not be blindly treated as the end goal. Especially static and retrospective benchmarks, in which no true unknown test set is employed, limit our ability to robustly validate a model's performance. Building on the collective expertise of a cross-industry consortium, we present a model validation framework consisting of five recommendations that would enable the community to move beyond aggregate metrics toward understanding where and why molecular property prediction models fail. We connect evaluation choices to real-world applications and case studies encountered in pharmaceutical research. The framework proposes splitting strategies that mimic realistic distribution shifts and expose common failure modes. We apply the recommended framework to a recently released dataset of absorption, distribution, metabolism, and excretion (ADME) properties. Across two complementary model algorithms, our case studies reveal four distinct failure modes (extrapolation, interpolation, representation, and evaluation), showing that model errors arise not only from distribution shift but also from limitations in molecular representations. Our results show that commonly used evaluation protocols can significantly overestimate performance and may not detect important model failure modes. All software and data are released via https://github.com/srijitseal/polaris.

19
Benchmarking Graph Neural Networks for Multi-Omics Cancer Subtyping using Methylation and Gene Expression Profiles

Schirmacher, J.; Maurer, M. C.; Metsch, J. M.; Ploesch, S.; Chereda, H.; Blumenthal, D. B.; Hauschild, A.-C.

2026-08-25 bioinformatics 10.64898/2026.08.21.745839 medRxiv
Top 0.1%
2.7%
Show abstract

Motivation: Graph Neural Networks (GNNs) have gained increasing interest in the biomedical domain, as the integration of prior knowledge and deep neural networks has the potential to enhance insights into molecular processes and disease mechanisms. However, a comprehensive and systematic assessment of model architectures, data modalities, graph structures, and their performance for graph signal classification in the biomedical domain is yet to be performed. In order to close this gap, we conducted a benchmarking study on multiple GNNs on a Protein-Protein Interaction (PPI) network for Kidney Renal Clear Cell Carcinoma and Breast cancer subtype prediction, performing an in-depth investigation of architectures, incorporating skip connections and various data modalities. Results: While none of the GNNs outperforms the structure-agnostic Multi-Layer Perceptron baseline, all of them can handle bimodal data (gene methylation and expression) and offer the ability to gain explainability based on PPIs. We offer practical guidelines for applying GNNs to graph signal processing tasks specifically for cancer classification. Depending on the underlying dataset and PPI structure employed, models on different data modalities outperform others. Overall, we suggest using ChebNet, which tends to outperform the Graph Convolutional Network and the Graph Attention Network in cancer subtype prediction. We recommend using GNN architectures that employ a simple flattening readout layer, as they provide better classification performance and faster training time than those with global average pooling. Additionally, we tested residual connections, but they had only an insignificant impact on classification performance.

20
A distribution-aware and functionally relevant novel framework for generation and discovery of bioactive peptides

Abhigyan, R.; Sood, V.; Arora, P.; Kaur, B.

2026-08-09 bioinformatics 10.64898/2026.08.04.742799 medRxiv
Top 0.1%
2.5%
Show abstract

Recent advances in artificial intelligence have accelerated the discovery of bioactive peptides by enabling computational exploration of the vast peptide sequence space. However, existing peptide generation approaches generally rely on either distribution-learning models, which generate biologically realistic sequences but do not consistently optimize functional activity, or optimization-based methods, which maximize prediction confidence while often deviating from the underlying distribution of experimentally validated peptides. To address this limitation, a two-phase generative-evolutionary framework is proposed that integrates distribution learning with evolutionary optimization. In the first phase, Variational Autoencoders (VAE), Autoregressive Transformers (ART), and Token Diffusion Transformers (TDT) are used to generate biologically plausible seed peptides. In the second phase, these peptides were used as initial seed for Hill Climbing optimization procedure that iteratively improves fitness function score. The proposed two-phase framework was evaluated using a dataset of experimentally validated IL-2-inducing peptides. Evaluation using independent IL-2 prediction models showed that Autoregressive Transformer combined with Hill Climbing achieved the best overall performance, achieving the mean IL-2 induction confidence score of 0.96 while reducing KL divergence from 2.26 for standalone Hill Climbing to 0.75. A case study on an independent IL-13 inducing peptide dataset showed similar trends, with ART initialized Hill Climbing achieving the mean IL-13 induction score of 0.99 while reducing KL divergence from 1.76 to 0.59. Overall, the framework provides a generalizable approach for balancing functional optimization and distributional realism and can be applied to peptide discovery and data augmentation in imbalanced biological datasets thereby generating high confidence peptides for wet lab validation. HighlightsO_LIProposed a two-phase framework for bioactive peptide generation with potential to address class imbalance in peptide classification tasks. C_LIO_LIPerformed a systematic comparison of distribution-learning and optimization-based approaches for peptide generation. C_LIO_LICombined distribution-learning models for sequence generation with optimization algorithms for improving peptide functional properties. C_LIO_LIDemonstrated the applicability of the proposed framework across multiple bioactive peptide datasets. C_LI